Subword-dependent speaker clustering for improved speech recognition

نویسندگان

  • Li Jiang
  • Xuedong Huang
چکیده

Speaker variability has a significant impact to the state-of-the-art speech recognition systems. Traditionally speaker clustering is performed without considering individual or class phonetic similarities across different speakers. In fact, clustered speaker groups may have very different degrees of variations for different phonetic classes. In this paper, speaker clustering is performed at subword level or subphonetic level. With one or more instances derived from clustering for each subword or subphonetic unit, we model speaker variation explicitly across different subword or subphonetic instances. In addition, we select from massive possible combinations of speaker-clustered subword models to form our initial model for speaker adaptation. Experiments show that subword-dependent speaker clustering is more effective than the traditional speaker clustering.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

Improved Bayesian Training for Context-Dependent Modeling in Continuous Persian Speech Recognition

Context-dependent modeling is a widely used technique for better phone modeling in continuous speech recognition. While different types of context-dependent models have been used, triphones have been known as the most effective ones. In this paper, a Maximum a Posteriori (MAP) estimation approach has been used to estimate the parameters of the untied triphone model set used in data-driven clust...

متن کامل

Speech Recognition Using Demi-Syllable Neural Prediction Model

The Neural Prediction Model is the speech recognition model based on pattern prediction by multilayer perceptrons. Its effectiveness was confirmed by the speaker-independent digit recognition experiments. This paper presents an improvement in the model and its application to large vocabulary speech recognition, based on subword units. The improvement involves an introduction of "backward predic...

متن کامل

Speech Recognition as Feature Extraction for Speaker Recognition

Information from speech recognition can be used in various ways in state-of-the-art speaker recognition systems. This includes the obvious use of recognized words to enable the use of text-dependent speaker modeling techniques when the words spoken are not given. Furthermore, it has been shown that the choice of words and phones itself can be a useful indicator of speaker identity. Also, recogn...

متن کامل

Constrained Subword Units for Speaker Recognition

Phonetic features have been proposed to overcome performance degradation in spectral speaker recognition in difficult acoustic conditions. The harmful effect of those conditions, however, is not restricted to spectral systems but also affects the performance of the open-loop phone recognisers on which phonetic systems are based. In automatic speech recognition, larger subword units and the use ...

متن کامل

Recent Progress in Robust Vocabulary-Independent Speech Recognition

This paper reports recent efforts to improve the performance of CMU's robust vocabulary-independent (VI) speech recognition systems on the DARPA speaker-independent resource management task. The improvements are evaluated on 320 sentences that randomly selected from the DARPA June 88, February 89 and October 89 test sets. Our first improvement involves more detailed acoustic modeling. We incorp...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:

دوره   شماره 

صفحات  -

تاریخ انتشار 2000